Papers with multi-modal Transformer
Visio-Linguistic Brain Encoding (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning. |
| Approach: | They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning . |
| Outcome: | The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing . |
Encoding and Controlling Global Semantics for Long-form Video Question Answering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods to find answers for long videos fail to reason over the whole sequence of video, leading to sub-optimal performance. |
| Approach: | They propose a state space layer to integrate global semantics into video . they use a gating unit to enable controllability over the flow of global semantic into visual representations. |
| Outcome: | The proposed framework is able to integrate global semantics into visual representations. |